Papers with Outcome Reward Models
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are prone to hallucination, especially during multihop tasks. |
| Approach: | They propose a hierarchical, erroraware discriminative PRM that classifies math errors at each step and combines finegrained signals to estimate step correctness. |
| Outcome: | The proposed model outperforms the prior best in a new stateof-theart PRMScore of 67.7 on a 400Ksample dataset . |
Logical Reasoning with Outcome Reward Models for Test-Time Scaling (2025.emnlp-main)
Copied to clipboard
| Challenge: | Logical reasoning is a critical benchmark for evaluating the capabilities of large language models (LLMs), but it is under-explored in deductive reasoning. |
| Approach: | They propose to use Chain-of-Thought to generate data using single and multiple samples to train ORMs. |
| Outcome: | The proposed model expands the type of errors covered in the training dataset, covering previously unexplored error types. |